Back

Journal of Chemical Information and Modeling

American Chemical Society (ACS)

Preprints posted in the last 90 days, ranked by how well they match Journal of Chemical Information and Modeling's content profile, based on 238 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit.

1
Benchmarking Docking Protocols for GPCR Allosteric Modulators

Thompson, T. D.; Miao, Y.

2026-08-20 bioinformatics 10.64898/2026.08.12.744492 medRxiv
Top 0.1%
71.0%
Show abstract

G protein-coupled receptor (GPCR) allosteric modulators (AMs) offer significant therapeutic advantages over orthosteric drugs, yet structure-based virtual screening lacks validated protocols accounting for the conformational complexity of GPCR allosteric sites. We benchmark docking protocols using PDB experimental structures and structural ensembles derived from Gaussian accelerated Molecular Dynamics (GaMD) simulations across four Class A GPCRs (including the muscarinic M2 and M4 receptors, the {beta}2-adrenergic receptor, and the C-C chemokine receptor type 2) with four programs (Glide HTVS, AutoDock Vina, DOCK3.8, and Boltz-2) against experimentally validated modulator libraries and property-matched decoys. GaMD ensemble docking improved early AM enrichment across all four targets under at least one program. Glide ensemble docking was the only protocol to consistently improve early AM recovery across all four targets, ranking known actives almost exclusively within the top 0.5% of compounds at CCR2 and improving M2R active recovery nearly 9-fold relative to the PDB structure. GaMD free-energy landscape topology governed ensemble re-ranking strategy selection: population-skewed landscapes favored top binding energy ranking (BEmin) while flat, multi-populated landscapes favored average binding energy ranking (BEavg), and at targets with dominant low-energy states, a single GaMD cluster matched or exceeded full ensemble or PDB performance. Taking the union of top percentile hits identified by both ensemble re-ranking methods, BEmin / BEavg, maximizes chemical diversity at the earliest percentiles. Program-specific scaffold recovery biases further motivated a consensus BEmin / BEavg approach to maximize hit diversity. The Boltz-2 deep-learning program showed minimal sensitivity to GaMD templates and underperformed conventional docking, suggesting its affinity predictions complement rather than replace physics- and empirical-based docking approaches for GPCR AM screening.

2
CHARMM-GUI Covalent Ligand Docker as a Web-based Molecular Docking Platform for Covalent Ligands

Kong, L.; Suh, D.; Im, W.

2026-07-16 biophysics 10.64898/2026.07.13.738313 medRxiv
Top 0.1%
62.8%
Show abstract

Covalent inhibitor research is an emerging topic in drug discovery due to its superior performance in specificity and inhibition effects. While molecular docking is a popular strategy in prediction and assessment of ligand conformations or poses in receptor proteins, covalent ligand docking requires nontrivial preparation efforts, as the ligand structure changes during the covalent complex formation. In order to facilitate molecular docking for covalent ligands, we have developed CHARMM-GUI Covalent Ligand Docker (CGUI-CLD), a new module for covalent ligand docking supported by AutoDock4. CGUI-CLD automates ligand preparation, supports ligand modification, implements docking simulation, and presents results through an intuitive user interface. A knowledge-based library built in CGUI-CLD currently supports 66 warheads and 8 amino acids, which can be used to automate the covalent ligand transformation from a pre-reaction to a post-reaction adduct form seamlessly. Moreover, CHARMM-GUI High-Throughput Simulator is integrated for rapid generation of multiple molecular dynamics simulation systems. CGUI-CLD is expected to significantly reduce a massive workload of covalent ligand docking and advance covalent ligand research. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/738313v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@1ccfefforg.highwire.dtl.DTLVardef@1794047org.highwire.dtl.DTLVardef@16b0368org.highwire.dtl.DTLVardef@acbc3e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOAbstract TOCC_FLOATNO C_FIG

3
Structural Context Determines Docking Engine Performance: A Family-Stratified Benchmark of Six Engines

Alejo, K.; Fisher, S.; Kalluri, T.; More, B.; Rajgure, H.; Panda, P. K.; Korban, C.; Chung, C.

2026-08-11 biochemistry 10.64898/2026.08.11.744016 medRxiv
Top 0.1%
60.5%
Show abstract

Molecular docking and co-folding engines are widely used to prioritize compounds for wet-lab validation, yet their accuracy is known to vary substantially across protein targets for reasons that remain only qualitatively understood. Here we benchmark six docking and co-folding engines (RevDock, DiffDock, Boltz2, AutoDock-GPU, rDock, and PandaDock) across 14 protein families, evaluating scoring power, ranking power, docking power, and physical validity. Rather than treating engine performance as protein-family-specific, we classify all 14 families into six mechanistic groups according to which of four scoring-function simplifications, rigid receptor, pairwise additivity, fixed point charges, and implicit solvent, is most severely stressed by that familys binding site. This framework helps explain, rather than simply describe, where each engine succeeds or fails: RevDocks CNN rescoring layer mitigates the pairwise additivity and fixed-charge limitations relative to physics-only scoring, achieving the highest overall pose accuracy (73.3% of poses [≤] 2.0 [A] RMSD), while Boltz2s sequence-based co-folding bypasses the rigid-receptor assumption and achieves comparable affinity correlation (mean Pearson r {approx} 0.60 for both engines). PandaDock, run with expanded conformational sampling, matches RevDock on pose accuracy (72.1% of poses [≤] 2.0 [A], lowest median RMSD at 0.96 [A]) and exceeds AutoDock-GPU on affinity correlation (mean r = 0.460), indicating that the performance of a physics-based scoring function is limited as much by search adequacy as by the scoring function itself. These results suggest that engine selection for a docking or co-folding campaign should be guided less by an engines aggregate benchmark ranking and more by which of these four structural and physical characteristics dominate the target of interest.

4
PAM-DB: Revealing Protein Activation Mechanisms for Next-Generation Rational Drug Discovery

Zhu, X.; Li, X.; Hou, Y.; Zhou, R.; Yan, Y.; Warshel, A.; Bai, C.

2026-08-22 biophysics 10.64898/2026.08.20.745895 medRxiv
Top 0.1%
60.2%
Show abstract

Current rational drug design relies predominantly on computational (CADD/AIDD) methods that model binding thermodynamics and static conformations of target proteins, primarily in their inactive states. However, the kinetic parameters that govern experimental efficacy-such as catalytic turnover and signaling potency-are determined by molecular interactions with transition states (TS), intermediate states (IS), and the entire continuum of conformations along the least free-energy activation pathway. The absence of this dynamic dimension has fundamentally limited the predictive power and success rate of conventional structure-based approaches. Here, we present a structural database that systematically maps the complete activation trajectories of pharmaceutically relevant targets, encompassing TS, IS, and all connecting conformational ensembles. This resource offers multiple strategic advantages for drug discovery: enabling rational targeting of previously "undruggable" proteins, facilitating biased agonism/antagonism design, revealing cryptic allosteric sites in inactive conformations, identifying novel transient pockets along the activation route, rationalizing the mechanisms of existing drugs, predicting mutational effects on activation barriers, and prospectively forecasting drug resistance and off-target liabilities. We demonstrate the utility of this database through representative case studies and provide implementation guidelines for integration into existing discovery pipelines. More detailed information can be found at our website: https://www.momedpamdb.com/en.

5
MolJam: A Multidimensional Framework for Assessing Molecular Dataset Quality and Its Impact on Machine Learning

Wang, P.; Shi, Z.; Gao, X.; Zhou, R.

2026-08-25 bioinformatics 10.64898/2026.08.21.746384 medRxiv
Top 0.1%
60.0%
Show abstract

High-quality molecular datasets are essential for reliable machine learning in cheminformatics and bioinformatics, yet dataset quality is rarely assessed systematically and its relationship with downstream model performance remains poorly understood. Here, we present MolJam, an open-source framework for quantitative assessment of molecular dataset quality across five dimensions-structural integrity, data quality, experimental information quality, chemical space coverage, and data distribution-using 12 standardized metrics. Application of MolJam to 11 MoleculeNet and eight ChEMBL-derived datasets revealed widespread and heterogeneous quality issues, including undefined stereochemistry in up to 70.72% of molecules, inconsistent molecular representations, and contradictory labels. We next asked whether improving these quality metrics necessarily improves machine learning performance. Refinement of the ESOL and Lipophilicity datasets increased their MolJam quality scores but produced mixed effects on predictive performance, suggesting a competing influence of reduced dataset size. Controlled ablation experiments further demonstrated that both dataset quality and data quantity contribute to model performance and, notably, that retaining molecules with incomplete stereochemical information can outperform their removal when the resulting gain in data quantity offsets the quality penalty. Thus, molecular dataset curation cannot be reduced to maximizing data cleanliness alone but requires balancing multiple dimensions of data quality against information loss. MolJam provides a standardized framework for diagnosing molecular dataset limitations, comparing benchmark quality, and quantitatively evaluating how data curation decisions influence downstream machine learning.

6
Overcoming the accuracy-generalization tradeoff in docking and scoring for prospective virtual screening

Petrosyan, G.; Altunyan, V.; Ghukasyan, T.; Abramyan, T. M.; Arakelov, G.; Davtyan, A.; Aghajanyan, T.; Nakipov, I.; Navasardyan, G.; Fahradyan, A.; Tunanyan, H.; Simonyan, A.; Janczyk, P. Ł; De Silva, D.; Saribekyan, H.; Tsidilkovski, L.; Arakelov, V.; Ginoyan, N.; Mnatsakanyan, H.; Ratnikov, M.; Smbatyan, K.; Papoyan, A.; Papoian, G. A.

2026-08-03 biophysics 10.64898/2026.08.03.742480 medRxiv
Top 0.1%
59.6%
Show abstract

Virtual screening promises access to tens of billions of synthetically accessible, diverse compounds, yet it is rarely used as a primary hit-discovery strategy in contemporary drug-discovery campaigns. We argue that this gap reflects the real-world underperformance of the underlying docking and scoring methods: classical docking is generalizable but limited in accuracy by simple functional forms and parsimonious parameterization, whereas recent machine-learning approaches are highly expressive but do not generalize well to novel molecules and pockets, their reported accuracy often inflated by train-test leakage. To address these challenges, we introduce DODock and DOScore, docking and scoring ML/physics hybrid frameworks that are also highly expressive, yet generalize much better out of distribution compared with the prior ML approaches. This generalization has been prospectively tested in several ways. First, DODocks blind prediction of a drug candidate binding to PCSK9 was compared to the crystal structure that was subsequently solved, recovering the binding pose to 1.2 [A] RMSD. We also used DODock and DOScore in prospective virtual screening campaigns against four therapeutic targets, spanning an ectoenzyme (CD73), a kinase (IRAK4), an extended-substrate protease (FXI), and an allosteric inhibition of protein-protein interface (IL17). These screens yielded many chemically novel, biochemically and cellularly active inhibitors. In the case of CD73, which is a historically challenging target for virtual screening, our screen resulted in a roughly hundredfold improvement in hit rate over a recent machine-learning screen. Our results indicate that the apparent ceiling in the accuracy of virtual screening that seemed to have somewhat plateaued in the last two decades is not fundamental, and that structure-based interrogation of ultralarge chemical space may eventually become a credible primary route to novel chemical matter.

7
Can SMILES be fragmented into a concatenable ordered sequence of retrosynthetically interesting string block ?

Reboul, E.; Prabakaran, H.; Baaden, M.; Waldispuhl, J.; Taly, A.

2026-08-26 bioinformatics 10.64898/2026.08.25.747180 medRxiv
Top 0.1%
57.8%
Show abstract

Molecules generated by deep learning models are often difficult to synthesize. Their synthetic accessibility can be improved with automated retrosynthetic analysis, which allows for identifying synthons. However, synthons in a SMILES can be scattered throughout the string depending on the path taken through the molecular graph used to generate the SMILES. We tested whether the ensemble of possible SMILES for a molecule can be used to generate a concatenable ordered sequence of string fragments (blocks) from SMILES that match potential synthons obtained through automated retrosynthetic analysis. We found that exhaustively sampling the SMILES space of a molecule improves the coverage of retrosynthetic breaks. We achieved full coverage of retrosynthetic bonds in string form for 85\% of the 1.9 million molecules in the MOSES dataset. Doing so allowed us to test our block SMILES in an unconditional de novo drug design test case with MolGPT and Monte Carlo Tree Search (MCTS). We found that using blocks as an LLM's token did degrade MolGPT performance due to the curse of dimensionality. However, using the SMILES selected by our blocking algorithm with the default SMILES tokenizer improved the reproduction of physico-chemical properties of samples and also improved uniqueness, novelty, and validity. The MCTS outperforms our MolGPT models in terms of validity and novelty. However, samples generated by the MCTS had physico-chemical properties that were further away from the MOSES baseline than the samples produced by molGPT, with an improved distribution of quantitative estimation of drug-likeness (QED).

8
Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design

Kumar, H.; Yang, Z.; Yu, Y.; Wen, J.; Kim, P.; Zhou, X.

2026-08-20 bioinformatics 10.64898/2026.08.14.744939 medRxiv
Top 0.1%
55.4%
Show abstract

Generative artificial intelligence is accelerating molecular design, yet the relative suitability of available models for different targets and stages of preclinical drug discovery remains unclear. Here we benchmarked 12 molecular generation and optimization methods across 176 curated protein-ligand systems spanning diverse therapeutic target classes, with experimentally validated ligands providing reference chemical space. The evaluated methods encompassed pocket-conditioned 3D generation, diffusion and flow-based modeling, autoregressive construction, reference-conditioned optimization and synthesis-aware design. Performance was assessed using operational robustness, chemical validity, uniqueness, molecular and scaffold diversity, quantitative estimate of drug-likeness, synthetic accessibility, docking, physicochemical and ADMET properties, and computational resource requirements. The results revealed architecture-dependent trade off such as receptor-conditioned methods exploited binding-pocket geometry, flow-based approaches enabled efficient sampling, reference-conditioned methods favored analogue generation, and synthesis-aware approaches improved chemical feasibility, but no method consistently optimized all criteria. To address the functional potential of generated molecules, we further developed a state-aware functional classifier (SAFC) that integrates molecular dynamics derived receptor ensembles, ensemble docking and protein ligand interaction graphs. SAFC provided dynamics-aware functional activity rankings for generated molecules that were partly complementary to docking, drug-likeness and synthetic accessibility scores. These findings support hybrid, stage specific deployment of generative models rather than reliance on any single architecture or evaluation metric. This study provides practical guidelines for generative AI based preclinical drug development processes.

9
A multi-agent molecular optimization framework leads to a rapid-recovery intravenous anesthetic candidate with an improved safety margin

Xue, Z.; Liu, X.

2026-08-20 bioinformatics 10.64898/2026.08.17.745149 medRxiv
Top 0.1%
54.0%
Show abstract

Lead optimization, the systematic refinement of therapeutic compounds through iterative structural modification, faces a dual challenge in modern drug discovery: navigating astronomically vast molecular design spaces while balancing conflicting demands on potency, pharmacokinetics, and safety. We present MASCOT (Multi-Agent SearCh for molecular OpTimization), a role-specialized multi-agent framework for molecular optimization. Integrated with a chemically constrained graph-editing search, MASCOT coordinates three specialized agents: a trade-off agent that reprioritizes competing objectives, a strategy agent that adapts how molecular edits are proposed, and a reflection agent that distills lessons from previous decisions. Computational experiments showed that MASCOT achieved the best performance over competing methods on six benchmark settings. On the SARS-CoV-2 main protease task, its mean docking-score improvement was 3.6 times that of the strongest baseline. Applied to the clinically used anesthetic remimazolam (RM), MASCOT prioritized RM-1, which showed a shorter liver microsomal half-life, higher brain exposure, and a larger therapeutic index than RM. Subsequent derivative design yielded RM-7. Extensive animal studies established RM-7 as a rapid-recovery intravenous anesthetic candidate with greater potency, faster functional recovery, a wider safety margin, and preserved flumazenil reversibility. These results demonstrate that multi-agent coordination can link adaptive molecular search to medicinal chemistry and experimental pharmacology.

10
StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization

Yang, B.; Xu, K.; Xiang, C.; Lee, B.; Xu, Y.; Li, T.; Shi, Y.; Sinitskiy, A.; Li, J.

2026-07-02 bioinformatics 10.64898/2026.06.28.735128 medRxiv
Top 0.1%
53.0%
Show abstract

Structure-based generative models (SBGMs) hold great promises for accelerating drug discovery by enabling target-aware molecular design. However, existing approaches face fundamental challenges: three-dimensional graph-based models can explicitly incorporate protein structural information but often generate chemically implausible molecules due to limited training data, while chemical language models (CLMs) produce chemically plausible molecules but struggle to effectively leverage three-dimensional structural information for structure-conditioned generation and hard to incorporate lead optimization functionality due to the nature of SMILES string. Here, we present StructureSAFE, a structure-aware chemical language model that resolves this trade-off by integrating protein structural and evolutionary encoders with the SAFE molecular representation via pretraining and finetuning training scheme, enabling both de novo hit identification and a comprehensive suite of lead optimization subtasks within a unified framework. Comprehensive benchmarking on the MolGenBench dataset demonstrates that StructureSAFE achieves state-of-the-art (SOTA) performance across multiple metrics, with particularly pronounced improvements in chemical plausibility relative to graph-based models lacking pretraining. Evaluation on a rigorously constructed held-out test set further confirms its ability to generate drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets on both hit identification and lead optimization setting. In silico case studies across four therapeutically relevant targets validate its capacity to generate chemically plausible molecules that recapitulate key binding interactions of known high-affinity ligands while proposing novel interactions for potential better affinity and exploring previously unknown regions of chemical space. Taking together, StructureSAFE represents a versatile and practical tool to provide high-quality candidate molecules for augmenting medicinal chemistry workflows in both hit identification and lead optimization campaigns.

11
Kinase inhibitors can change protonation or tautomeric state upon binding

Ranepura, G. A.; Chowdhury, S. I.; Rosenzweig, E. A.; Rustenburg, A. S.; Lopez-Rios de Castro, R.; Mao, J.; Chodera, J. D.; Singh, S.; Gunner, M. R.

2026-07-30 biophysics 10.64898/2026.07.27.741060 medRxiv
Top 0.1%
52.5%
Show abstract

The binding affinity of a ligand to a protein is influenced by the protonation and tautomeric states of both partners. However, this relationship remains under-investigated due to the limited availability of computational tools capable of considering all charge and tautomer states in a scalable manner to study clinically relevant systems. Here, we use Multi-Conformation Continuum Electrostatics (MCCE) to calculate the protonation and tautomer distributions of nine kinase domains bound to 18 FDA-approved inhibitors while considering their Boltzmann-ensemble. Our simulations show that protein net charge and proton distribution remain largely stable even upon binding charged inhibitors. Our results find that individual inhibitor charges are dynamic, frequently increasing, or decreasing upon binding a specific protein target. Kinase-inhibitor binding significantly shifts the relative probabilities of low-energy states ({Delta}G < 2.5 kcal/mol), though it does not recruit higher-energy conformers into the bound population. Our consideration of all possible charge states and tautomers enable us to identify when tautomer have significant significant free binding energy differentials (3-6kcal/mol). In turn, we find that minority species can become the dominant component in the bound state, emphasizing the necessity of considering ensemble-wide protonation and tautomer states to accurately predict protein-ligand binding energetics.

12
MolMAE: A Surface-Centric Multimodal Masked Autoencoder for Molecular Representation Learning

Li, J.

2026-07-14 bioinformatics 10.64898/2026.07.11.737987 medRxiv
Top 0.1%
50.1%
Show abstract

Molecular representation learning has become a central component of modern computational drug discovery. Existing molecular foundation models mainly rely on SMILES strings, two-dimensional molecular graphs, or three-dimensional atomic coordinates. However, many molecular properties are ultimately governed by the molecular surface, where intermolecular recognition, solvation, electrostatic complementarity, and ligand-protein interactions occur. In this work, we propose MolMAE, a surface-guided multimodal masked autoencoder for molecular representation learning. MolMAE takes molecular surface point clouds, three-dimensional molecular graphs, and SMILES-derived fragment and functional-group tokens as complementary input modalities, and learns a unified multimodal molecular embedding through functional-group-aligned masked autoencoding. During pretraining, chemically corresponding local regions are jointly masked across surface, graph, fragment, and functional-group views, forcing the model to reconstruct missing geometric, physicochemical, structural, and semantic information from the remaining context. While molecular surface reconstruction serves as the primary pretraining objective, graph-, fragment-, and functional-group-level reconstruction tasks provide complementary supervision that encourages the model to capture molecular topology, bonding patterns, stereochemistry, local chemical environments, and substructure organization. In addition to reconstructing surface geometry, MolMAE reconstructs surface-associated physicochemical fields, including electrostatic potential and Fukui-related descriptors, enabling the model to learn chemically meaningful surface representations. Pretrained on approximately 261K lead-like bioactive molecules, MolMAE achieves strong performance on the ESOL benchmark under scaffold splitting and competitive performance across multiple molecular property prediction tasks. These results suggest that molecular surface-guided pretraining can complement conventional graph-, sequence-, and atom-coordinate-based molecular representations, especially for property prediction tasks influenced by exposed surface geometry and surface-associated physicochemical patterns.

13
A multimodal representation learning platform for accurate molecular ADMET prediction

Luo, Z.; Huang, D.; Shao, Y.; Yu, Q.; Li, Y.

2026-08-25 bioinformatics 10.64898/2026.08.24.746660 medRxiv
Top 0.1%
44.8%
Show abstract

Accurate ADMET prediction is essential for prioritizing compounds before costly experimental validation, yet ADMET tasks are highly heterogeneous. Properties such as solubility, permeability, protein binding, clearance, transporter activity and toxicity are governed by different molecular signals, ranging from local functional groups and physicochemical descriptors to bonded topology and three-dimensional geometry. Consequently, a single molecular representation or backbone is unlikely to be optimal across all ADMET tasks. We present Trimole-Hybrid, a task-wise multimodal framework that addresses ADMET heterogeneity by selecting or combining predictors built from complementary molecular representations. Trimole-Hybrid constructs a candidate pool of SMILES-, graph-, geometry-sensitive EPT/3D- and chemical descriptor-based predictors. For each task, Trimole-Hybrid selects the best-performing predictor to obtain the final prediction. On 22 Therapeutics Data Commons ADMET benchmarks, Trimole-Hybrid exceeded the public TDC top-1 methods on 10 tasks and ranked within the top 10 for 21 tasks. Ablation studies confirmed the contribution of both complementary multimodal molecular representations and task-specific ensemble strategies. In two small-molecule case studies, Trimole-Hybrid shows sensitivity to changes in essential functional motifs, suggesting its ability to capture ADMET-relevant molecular substructures.

14
CARS: A General Force Field for Carotenoids

Nikolaev, A.; Orlov, Y.; Khanina, V.; Gushchin, I.

2026-08-05 bioinformatics 10.64898/2026.07.31.742033 medRxiv
Top 0.1%
40.9%
Show abstract

Carotenoids are structurally diverse isoprenoid pigments that play central roles in photosynthesis, photoprotection, membrane organization, and cellular signaling. Despite their biological and technological importance, atomistic simulations of carotenoids remain limited by the lack of a transferable force field spanning the chemical diversity of naturally occurring compounds, including glycosylated and acylated derivatives. Here we present CARS (CARotenoidS), a transferable force field for carotenoids that integrates seamlessly with the AMBER family of biomolecular force fields. Parameters were systematically optimized against 22 957 r2SCAN-3c reference energies for 25 representative molecular fragments, yielding an accurate description of polyene conformational energetics, ring rotations, and molecular geometries. Across diverse validation sets, CARS substantially outperforms GAFF2 and provides improved agreement with quantum-chemical reference data for glycosylated and acylated carotenoids. For zeaxanthin, CARS also surpasses OPLS-AA, CGenFF, and previously published carotenoid-specific parameters set in reproducing conformational energetics and structural properties. Two complementary parameter sets are provided: CARS for glycosylated and non-lipidated carotenoids, and CARS+Lipid21 for carotenoids containing saturated or monounsaturated lipid chains. By providing the first unified and transferable parameterization covering the structural diversity of natural carotenoids while remaining fully compatible with established AMBER force fields, CARS removes the need for molecule-specific reparameterization and enables reliable molecular simulations of carotenoids in proteins, membranes, and other complex biological assemblies.

15
Toward Robust Characterization of Dynamic Binding Pockets: Lessons from the HBV Capsid Assembly Modulator Site

Perez-Segura, C.; Scott, L. W.; Zlotnick, A.; Hadden-Perilla, J. A.

2026-08-10 biophysics 10.64898/2026.08.06.743403 medRxiv
Top 0.1%
40.6%
Show abstract

Protein function often depends on ligand binding pockets that fluctuate among conformational states, altering their size, shape, topology, and accessibility, yet quantitative comparison of these dynamic cavities remains challenging because their boundaries are often inherently ambiguous. The measure volinterior algorithm uses fuzzy-boundary detection to characterize enclosed molecular spaces; here, the hepatitis B virus (HBV) capsid assembly modulator (CAM) binding site is used as a model system to develop and validate a practical workflow for applying the method to dynamic protein binding pockets. The resulting methodology provides practical guidance for parameter selection and evaluation, establishes a standardized protocol for quantitative characterization of the HBV CAM pocket, and demonstrates robust, reproducible performance across conformational ensembles derived from molecular dynamics (MD) simulations. More broadly, this work provides a reproducible strategy for adapting measure volinterior to other dynamic binding pockets, enabling consistent comparison of pocket geometry among independent structural studies. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/743403v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@d460b5org.highwire.dtl.DTLVardef@11915c2org.highwire.dtl.DTLVardef@1e37524org.highwire.dtl.DTLVardef@1fc1b7_HPS_FORMAT_FIGEXP M_FIG C_FIG

16
PG-MLD: Physics-Guided Molecular Representation Learning via Dynamic 3D Trajectory Distillation

liu, Z.; Wu, Z.; Chen, Z.; Gao, X.; Yu, B.

2026-08-02 bioinformatics 10.64898/2026.07.29.741404 medRxiv
Top 0.1%
39.6%
Show abstract

Molecular representation learning underpins molecular property prediction and drug design by capturing molecular structure-property relationships. SMILES-based molecular language models learn chemical semantics from large-scale unlabeled data and support efficient inference. However, the one-dimensional nature of SMILES constrains their ability to capture 3D geometry and conformational evolution, whereas 3D molecular models require conformer generation and substantial computational resources. To bridge this gap, we propose PG-MLD, a dynamic 3D-to-1D physical knowledge distillation paradigm for molecular representation learning. PG-MLD constructs a dynamic 3D physical teacher by combining equivariant geometric encoding with Liquid Time-Constant modeling to capture 3D geometry, atom-level electronic descriptors, and conformational evolution. PG-MLD subsequently distills the learned trajectory knowledge into SMILES-based students through atom- and molecule-level representation alignment and cross-modal contrastive learning, with masked language modeling retained where supported. The distilled students perform downstream tasks using only SMILES, without conformer generation or molecular dynamics simulations. Experiments on MoleculeNet show that PG-MLD improves overall property prediction performance across three molecular language student architectures while maintaining SMILES-only inference. The learned representations also encode 3D geometry and conformational dynamics more effectively, demonstrating that dynamic 3D physical knowledge can be transferred to SMILES-based molecular language models with different architectures.

17
Computationally mapping olfactory receptors to odor percepts using docking energy scores

Guan, Y.

2026-07-05 bioinformatics 10.64898/2026.06.30.735315 medRxiv
Top 0.1%
39.5%
Show abstract

Mapping the olfactory factors directly to perceived smells has broad implications for neuroscience, chemistry, and medicine. Since the discovery of olfactory genes in 1991, the completion of the olfactory code has been hindered by two obstacles: the unknown combinatorial principles by which ~400 receptors collectively encode thousands of perceptually distinct odors, and the absence of a complete functional map from the olfactory receptors to conscious perception. We investigated this problem by using the binding energy scores of odorants with the olfactory receptors. We first showed that using only docking scores, we can predict the smell percepts with an accuracy similar to a full set of chemical fingerprints, combining the two resulted in even better performance. This supports there is a direct relationship between olfactory receptors and specific smells. Next, we turned on the olfactory receptors one by one by iterative training and simulation, and produced corresponding specific perception profiles for each olfactory receptor. The generated matrix is sparse, with only 1-2 smell types activated for each olfactory receptors. Despite the limitation of the size of the training data, it suggests the possibility of a low-dimensional combinatorial principle underlying thousands of smells that humans can perceive. We confirmed the prediction by a list of well-known olfactory receptors. We applied the model trained on single chemicals to mixtures, confirming competitive binding was the driving force for smell specificity. This strong performance, surprisingly, is established on the simplest modeling of the binding scores on hundreds of chemicals, and we believe the mapping can become more accurate if more complicated structural modeling techniques and more data are used.

18
Rank-Resolved Multi-Engine Docking and Optuna-Optimized Re-Ranking with ProDock for Virtual Screening

Le, L. H. S.; Pham, T.-A.; Tran, N.-T. N.; Van-Nguyen, P.-C.; Phan, T. L.; Truong, T. N.

2026-07-24 pharmacology and toxicology 10.64898/2026.07.24.740457 medRxiv
Top 0.1%
39.5%
Show abstract

False positives in virtual screening often arise when a single docking score or top-ranked pose is treated as sufficient evidence for binding. We extend the previously introduced ProDock software from a database-backed docking platform into a rank-resolved, multi-engine workflow for automated preparation, docking, pose analysis, and optimized re-ranking. The extended workflow combines local docking with GNINA and global docking with DiffDock with pose-level descriptors, namely binding-site occupancy, ligand localization, interaction-fingerprint similarity, and steric clash counts, together with Optuna -based threshold optimization. Across 43 DUDE-Z targets, the archived benchmark outputs reported higher enrichment values for CNN-based GNINA scores after optimization. CNNaffinity PR-AUC changed from 0.197 to 0.294 and LogAUC from 0.708 to 0.763, whereas empirical affinity ROC-AUC changed from 0.770 to 0.758. Structural investigation of re-docked actives showed that re-ranked poses were more native-like, with improved binding-site occupancy, reduced centroid displacement, and greater recovery of co-crystal interactions. The extension provides a reproducible framework for combining complementary docking engines with interpretable pose-level metrics before hit selection, thereby aiding the identification of true-positive candidates in virtual screening.

19
Mavchen 1: A Conformational Ensemble Platform for Protein Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

Varghese, R.; Tiwary, P.; Oswal, K.

2026-07-29 bioinformatics 10.64898/2026.07.26.740840 medRxiv
Top 0.1%
39.3%
Show abstract

Deep learning structure predictors, most prominently AlphaFold2 (the field-standard tool benchmarked against throughout this study), have substantially expanded access to protein structural information, yet characteristically return a single static conformation per target. This is an incomplete representation of the binding-competent state for the many pharmacologically relevant targets whose recognition geometry is intrinsically dependent on receptor flexibility, including cryptic-pocket, induced-fit, and water-mediated binding mechanisms. We present a category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein-ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes. Considering the most accurate pose available from each methods full candidate output, ensemble-derived poses achieved lower RMSD to the experimental structure than AlphaFold on 21 of 29 targets (72.4%), with a mean RMSD of 3.39 [A] versus 5.60 [A]: a clear, statistically decisive advantage (paired Wilcoxon signed-rank test, W = 93.0, p = 0.0060). Rather than being diffuse, this advantage was concentrated precisely where mechanistic theory predicts it should be: in induced-fit and water-mediated categories, the classes in which static-structure prediction is expected to be least representative of the bound state: a result that constitutes direct, quantitative confirmation of the ensemble hypothesis, not merely a favorable average. Independent assessment against a field-standard physical-validity framework confirmed that this accuracy gain was achieved without any trade-off in chemical or geometric realism. We further quantify, rather than assume, the extent to which this advantage is recoverable by fully autonomous pose selection, using a proprietary ensemble-aware scoring model with no access to the correct answer, and report a substantial, discriminative signal (cross-validated mean AUC 0.92) with a partial, and clearly characterized, recovery under the strictest accuracy criteria (mean AUPR 0.36), which we identify as the principal, now precisely quantified, determinant of near-term translational progress. Under this same fully autonomous, ground-truth-blind setting, AlphaFolds own top-ranked poses currently match or modestly exceed Mavchen-1s autonomously selected poses on strict success-rate criteria (e.g., 17.2% vs. 20.7% at the combined RMSD-and-validity threshold), a result we report without qualification as the clearest current benchmark for near-term development. Together, these results provide compelling, statistically rigorous evidence that conformational ensemble sampling is a mechanistically grounded and substantial source of improved pose accuracy relative to static-structure prediction, and establish a quantitative benchmark against which continued methodological development can be measured and demonstrably improved upon.

20
kontakteUR: transforming coordinates to chemical intuition to focus on essential interactions in biomolecular systems

Scherlo, M.; Wippermann, E.; Fuertges, T.; Kuenne, R.; Yelboga, A.; Ruetten, F.; Boeckmann, M.; Hoeweler, U.; Rudack, T.

2026-06-22 biochemistry 10.64898/2026.06.19.732925 medRxiv
Top 0.1%
39.3%
Show abstract

Molecular interactions govern cellular function, making them essential to discover biomolecular mechanisms by unravelling structure-function relationships. The rapid growth of AI-based prediction, experimental determination, and molecular dynamics simulations generates structural data at an unprecedented scale. However, structural information is typically represented as Cartesian coordinates, leaving chemical interactions and conformational relationships largely implicit. We introduce a high-throughput framework transforming structural geometry into a standardized, compact contact space. Moving beyond simple distance cutoffs, it provides a chemically and geometrically informed representation of various residue-residue interactions, their temporal changes, and conformations at residue-level resolution. Our contact-space representation enables systematic comparison and classification even for large-scale analysis. Case studies spanning structure comparison or studies of protein-protein, protein-ligand, protein-RNA, and antibody-antigen complexes, demonstrate how contact-space analysis reveals interaction patterns, identifies key mutation sites, and links structural features to experimental observations. With these and further applications, kontakteUR elucidates biomolecular function and assists targeted protein design, with results suited for further processing by artificial intelligence algorithms.